Skip to content

feat(model/gemini): resume generations paused with a continuation token - #1715

Open
baptmont wants to merge 1 commit into
mainfrom
baptmont/gemini-continuation
Open

baptmont wants to merge 1 commit into
mainfrom
baptmont/gemini-continuation

Conversation

@baptmont

@baptmont baptmont commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Linked issue

No issue yet. This implements the follow-up noted in #1714: ADK could see a CONTINUATION finish but could not resume from it.

Problem: a Gemini model with decoding continuation stops a generation that reaches its per-request output limit with finish reason CONTINUATION and an opaque Candidate.ContinuationToken. ADK Go returned that partial output as the final answer, so long answers were cut off.

Solution: the Gemini model resumes the generation itself, as ADK Java (google/adk-java#1600) and ADK Kotlin (google/adk-kotlin#615) do:

  • It resends the original contents plus the output so far as a model turn, with the token in GenerateContentConfig.ContinuationToken, until the generation finishes.
  • It returns one response with the joined content and the usage summed across requests (counts summed, per-modality counts summed by modality). The finish reason and any error code are the last request's, so a resumed generation that ends blocked still reports why.
  • Streaming: the resumed requests feed the same aggregator, so the caller sees one continuous stream. A paused chunk's finish reason is cleared only when the generation will be resumed, so the turn is not reported complete early.
  • It stops resuming, and returns what it has with finish reason CONTINUATION, after 256 resumes or when the same token comes back twice (no progress). MAX_TOKENS is not resumed, since the API applies maxOutputTokens to the whole generation. A CONTINUATION stop with no token logs a warning and returns the partial output.
  • Parts are copied before the aggregator and the flow see them, so a resend carries what the model sent, not, for example, the client function-call IDs the flow sets later.
  • No new exported API.

Two deliberate differences from Java and Kotlin, also stated in the code comments:

  1. Resends retry transient failures (408, 429, 5xx, transport errors; 8 attempts, 1s to 60s backoff), because failing a resend would discard the output so far. This does not apply when the request or the client sets HTTPOptions.RetryOptions. The first request of a generation is not retried, as before.
  2. A part with empty text and only a thought flag or signature is not resent, even with a signature. genai omits empty text, so such a part encodes with no data and the API rejects the resend with 400 ... required oneof field 'data'. I hit this live on a streamed resend: Gemini ends a stream with such a part. ADK Java keeps the signed ones because it sends "text": "".

Behavior change

  • Gemini model, generation stopped with CONTINUATION and a token: before, the caller got the first request's partial output with finish reason CONTINUATION. Now ADK sends more requests and returns the whole generation, with summed usage and the final finish reason. One GenerateContent call can now make up to 257 model requests, and its latency and token use grow to match.
  • Resends retry transient failures as described above. Nothing else about retries changes.
  • Nothing changes for models or responses that never stop with CONTINUATION.

Testing Plan

Unit Tests:

  • All unit tests pass locally.

Module loop green: go build -mod=readonly work, go test -race -mod=readonly -count=1 -shuffle=on work, golangci-lint run (v2.3.1, 0 issues) and go mod tidy -diff (no output) in both modules.

New tests in model/gemini/continuation_test.go, all against an httptest fake of the Gemini API:

  • TestContinuation_Generate: resumed until finished (resent contents and summed usage checked), same token twice, MAX_TOKENS not resumed, paused without a token.
  • TestContinuation_GenerateStream, TestContinuation_GenerateStreamGivesUp: partials, turn completion only at the real end, and CONTINUATION when the model gives up.
  • TestContinuation_ResendCopiesParts, TestContinuation_ResendDropsEmptyParts, TestContinuation_BlockedAfterResume, TestContinuation_StopsAtMaxResumes.
  • TestContinuation_RetriesResumingRequest: resend retried, first request not retried, giving up after 8 attempts, client's and caller's retry options respected.
  • TestAppendParts, TestAddUsage.

With your source change reverted and your tests kept, which test fails?

With gemini.go reverted: TestContinuation_Generate, TestContinuation_GenerateStream, TestContinuation_GenerateStreamGivesUp, TestContinuation_ResendCopiesParts, TestContinuation_ResendDropsEmptyParts, TestContinuation_StopsAtMaxResumes and TestContinuation_RetriesResumingRequest. I also disabled each guard one at a time: the max-resumes bound, the repeated-token check, the empty-part drop, thought/answer separation when joining text, first-signature-wins, summing each usage field and the per-modality counts, converting before overlaying the joined content, and the retry conditions. The suite failed every time.

Manual End-to-End (E2E) Tests:

Run live through ADK's Gemini model against a pre-release Gemini model with decoding continuation, configured to pause every 300 decoding steps. The prompt asks for 100 numbered lines with computed values.

  • Non-streaming: 15 resumes, all 100 lines present and correct, usage summed across the 16 requests.
  • Streaming: an earlier revision failed the first resend with the 400 described above. That led to the empty-part drop. With the fix it passes: 15 resumes, and all 100 lines arrive once and in order.

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules. (genai v1.72.0, chore(deps): bump google.golang.org/genai from 1.71.0 to 1.72.0 #1714)

Additional context

The adk-python port of the same behavior is not public yet, so this PR cites Java and Kotlin for parity:

  • ADK Java: core/src/main/java/com/google/adk/models/GeminiContinuation.java (MAX_RESUMES at line 42, the same-token stop at lines 103-112, the resend's model turn at lines 149-154, summed usage starting at line 202) and the stream-terminator check in Gemini.java:436-443.
  • ADK Kotlin: core/src/commonMain/kotlin/com/google/adk/kt/models/GeminiContinuation.kt.

Known follow-up shared with Java and Kotlin, not addressed here: when function-call arguments are streamed (PartialArgs), each chunk's part is recorded as received, so a paused stream resends them as separate parts.

A Gemini model with decoding continuation stops a generation that reaches
its per-request output limit with finish reason CONTINUATION and an opaque
continuation token. Until now ADK Go returned that partial output as the
final answer.

The Gemini model now resends the request with the output so far appended
as a model turn and the token in GenerateContentConfig.ContinuationToken,
until the generation finishes. It returns one response with the joined
content and the usage summed across requests; the finish reason and any
error code are the last request's. In streaming mode the resumed requests
feed the same aggregator, and a paused chunk's finish reason is cleared
only when the generation will be resumed. This matches ADK Java
(google/adk-java#1600) and ADK Kotlin (google/adk-kotlin#615).

The model stops resuming, and returns what it has with finish reason
CONTINUATION, after 256 resumes or when the same token comes back twice.
MAX_TOKENS is not resumed, since the API applies maxOutputTokens to the
whole generation.

Two differences from ADK Java and Kotlin:
- A resend retries transient failures (8 attempts) unless the request or
  the client sets retry options, because failing it would discard the
  output so far. The first request is not retried, as before.
- A part with empty text and only a thought flag or signature is not
  resent, even with a signature. genai omits empty text, so the part
  would carry no data and the API rejects it with a 400; ADK Java keeps
  signed ones because it sends the empty text.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant