Conversation
A Gemini model with decoding continuation stops a generation that reaches its per-request output limit with finish reason CONTINUATION and an opaque continuation token. Until now ADK Go returned that partial output as the final answer. The Gemini model now resends the request with the output so far appended as a model turn and the token in GenerateContentConfig.ContinuationToken, until the generation finishes. It returns one response with the joined content and the usage summed across requests; the finish reason and any error code are the last request's. In streaming mode the resumed requests feed the same aggregator, and a paused chunk's finish reason is cleared only when the generation will be resumed. This matches ADK Java (google/adk-java#1600) and ADK Kotlin (google/adk-kotlin#615). The model stops resuming, and returns what it has with finish reason CONTINUATION, after 256 resumes or when the same token comes back twice. MAX_TOKENS is not resumed, since the API applies maxOutputTokens to the whole generation. Two differences from ADK Java and Kotlin: - A resend retries transient failures (8 attempts) unless the request or the client sets retry options, because failing it would discard the output so far. The first request is not retried, as before. - A part with empty text and only a thought flag or signature is not resent, even with a signature. genai omits empty text, so the part would carry no data and the API rejects it with a 400; ADK Java keeps signed ones because it sends the empty text.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Linked issue
No issue yet. This implements the follow-up noted in #1714: ADK could see a
CONTINUATIONfinish but could not resume from it.Problem: a Gemini model with decoding continuation stops a generation that reaches its per-request output limit with finish reason
CONTINUATIONand an opaqueCandidate.ContinuationToken. ADK Go returned that partial output as the final answer, so long answers were cut off.Solution: the Gemini model resumes the generation itself, as ADK Java (google/adk-java#1600) and ADK Kotlin (google/adk-kotlin#615) do:
GenerateContentConfig.ContinuationToken, until the generation finishes.CONTINUATION, after 256 resumes or when the same token comes back twice (no progress).MAX_TOKENSis not resumed, since the API appliesmaxOutputTokensto the whole generation. ACONTINUATIONstop with no token logs a warning and returns the partial output.Two deliberate differences from Java and Kotlin, also stated in the code comments:
HTTPOptions.RetryOptions. The first request of a generation is not retried, as before.400 ... required oneof field 'data'. I hit this live on a streamed resend: Gemini ends a stream with such a part. ADK Java keeps the signed ones because it sends"text": "".Behavior change
CONTINUATIONand a token: before, the caller got the first request's partial output with finish reasonCONTINUATION. Now ADK sends more requests and returns the whole generation, with summed usage and the final finish reason. OneGenerateContentcall can now make up to 257 model requests, and its latency and token use grow to match.CONTINUATION.Testing Plan
Unit Tests:
Module loop green:
go build -mod=readonly work,go test -race -mod=readonly -count=1 -shuffle=on work,golangci-lint run(v2.3.1, 0 issues) andgo mod tidy -diff(no output) in both modules.New tests in
model/gemini/continuation_test.go, all against anhttptestfake of the Gemini API:TestContinuation_Generate: resumed until finished (resent contents and summed usage checked), same token twice,MAX_TOKENSnot resumed, paused without a token.TestContinuation_GenerateStream,TestContinuation_GenerateStreamGivesUp: partials, turn completion only at the real end, andCONTINUATIONwhen the model gives up.TestContinuation_ResendCopiesParts,TestContinuation_ResendDropsEmptyParts,TestContinuation_BlockedAfterResume,TestContinuation_StopsAtMaxResumes.TestContinuation_RetriesResumingRequest: resend retried, first request not retried, giving up after 8 attempts, client's and caller's retry options respected.TestAppendParts,TestAddUsage.With your source change reverted and your tests kept, which test fails?
With
gemini.goreverted:TestContinuation_Generate,TestContinuation_GenerateStream,TestContinuation_GenerateStreamGivesUp,TestContinuation_ResendCopiesParts,TestContinuation_ResendDropsEmptyParts,TestContinuation_StopsAtMaxResumesandTestContinuation_RetriesResumingRequest. I also disabled each guard one at a time: the max-resumes bound, the repeated-token check, the empty-part drop, thought/answer separation when joining text, first-signature-wins, summing each usage field and the per-modality counts, converting before overlaying the joined content, and the retry conditions. The suite failed every time.Manual End-to-End (E2E) Tests:
Run live through ADK's Gemini model against a pre-release Gemini model with decoding continuation, configured to pause every 300 decoding steps. The prompt asks for 100 numbered lines with computed values.
Checklist
Additional context
The adk-python port of the same behavior is not public yet, so this PR cites Java and Kotlin for parity:
core/src/main/java/com/google/adk/models/GeminiContinuation.java(MAX_RESUMESat line 42, the same-token stop at lines 103-112, the resend's model turn at lines 149-154, summed usage starting at line 202) and the stream-terminator check inGemini.java:436-443.core/src/commonMain/kotlin/com/google/adk/kt/models/GeminiContinuation.kt.Known follow-up shared with Java and Kotlin, not addressed here: when function-call arguments are streamed (
PartialArgs), each chunk's part is recorded as received, so a paused stream resends them as separate parts.