Skip to content

feat(telemetry)!: default to experimental GenAI semconv, with ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN=otel_semconv_1_36 rollback - #1635

Open
RKest wants to merge 1 commit into
google:mainfrom
RKest:feat/telemetry-semconv-default
Open

RKest wants to merge 1 commit into
google:mainfrom
RKest:feat/telemetry-semconv-default

Conversation

@RKest

@RKest RKest commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Please ensure you have read the contribution guide before creating a pull request.

Linked issue

Problem:
ADK Go emits three telemetry formats side by side:

  • stable-semconv (v1.36) log events;
  • experimental-semconv content on generate_content;
  • ADK-specific gcp.vertex.agent.tool_call_args / tool_response attributes, which record tool arguments and results on every execute_tool span whatever the content-capture setting.

The OpenTelemetry GenAI instrumentations now default to the experimental conventions, and adk-python is migrating the same way behind ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN (_schema_version.py#L72).

Solution:
The experimental GenAI semantic conventions (as of semantic conventions v1.44.0) become the default. ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN=otel_semconv_1_36 is a full rollback to today's format, and otel_semconv_1_44 names the default. Values are case-insensitive, and adk-python's 1 and 2 are accepted for them.

Default (otel_semconv_1_44) otel_semconv_1_36 (legacy)
Model-call logs One gen_ai.client.inference.operation.details per call. Finish reasons, usage and gen_ai.provider.name are always present; messages only with EVENT_ONLY / SPAN_AND_EVENT (or true). gen_ai.system.message, gen_ai.user.message, gen_ai.choice
Message content on spans (SPAN_ONLY / SPAN_AND_EVENT) Structured values (attribute.SliceValue / MapValue), as semconv asks when the SDK supports them JSON strings
execute_tool payloads gen_ai.tool.call.arguments / gen_ai.tool.call.result (the result only when the call succeeded), only with SPAN_ONLY / SPAN_AND_EVENT gcp.vertex.agent.tool_call_args / tool_response, always
execute_tool (merged) No payload attributes; each call's own span carries them. tool_call_args: "N/A" plus the merged event
Root span invoke_workflow {root} above a non-workflow root agent, as adk-python's record_invocation does invoke_agent {root}
Nested invoke_workflow (e.g. agent-as-tool) gen_ai.workflow.nested: true, as in node_tracing.py#L223 unchanged

Structure:

  • internal/telemetry/instrumentation.go is the only way to trace a model call or a runner invocation. generateContent wraps its model call in InstrumentGenerateContent, which owns the generate_content span, the legacy log events and the new event, so the span and its logs cannot drift apart. runner.Run wraps its body in InstrumentInvocation. The pieces they compose are now unexported.
  • The schema version is read from the environment on each use, as adk-python's resolve_schema_version does. There is no package-level state.
  • Span size guard: a content attribute whose JSON encoding exceeds 60 KiB is left off the span, in both schemas. I checked telemetry.googleapis.com, the endpoint ADK exports spans to: it rejects the whole export request when one span attribute is over 64 KiB, and it measures a structured value by its encoding (a 90 KiB SliceValue made of 30 KiB strings was rejected). Events are not capped by ADK; the log SDK's OTEL_LOGRECORD_ATTRIBUTE_VALUE_LENGTH_LIMIT applies to each string inside them.

Other changes:

  • adk-web: debug log records now include their attributes, which is where the bundled UI reads the new event from. The debug server rewrites message parts the UI's schema rejects (reasoning, uri, blob content, non-object tool payloads) into shapes it accepts. Without that, one thought part blanks the whole Traces panel.
  • Tests: test(telemetry): replace hand-written telemetry expectations with record/replay goldens #1640 already records every case under both schema values. This PR changes 32 of the 40 otel_semconv_1_44 goldens and none of the otel_semconv_1_36 ones. The other 8 are workflow failures that end before any model call.

Deliberate divergences from adk-python (also commented in code):

  • Default. adk-python still defaults to schema 1 outside Agent Engine, and gates its experimental events on OTEL_SEMCONV_STABILITY_OPT_IN. Go goes straight to the end state of that migration plan, with one env var.
  • Tool attributes. adk-python's schema v2 has not yet moved execute_tool to gen_ai.tool.call.* (step 3 of its plan). Go does it now, because the legacy attributes capture content by default.
  • Structured span content. adk-python records span content as JSON strings (_experimental_semconv.py#L724), because its OpenTelemetry API has no structured span attributes. otel-go has had them since v1.44 (SLICE) and v1.45 (MAP).
  • Streaming. adk-python records one output message per streamed chunk. Go records the final aggregated response, since semconv says one choice is not split across messages.

Behavior change

Breaking change for telemetry consumers. After upgrading:

  • The legacy log events are replaced by gen_ai.client.inference.operation.details.
  • Message content on spans is a structured value rather than a JSON string. Cloud Trace shows it as JSON either way.
  • execute_tool no longer records tool arguments or results unless span content capture is on, and then under gen_ai.tool.call.*.
  • An LLM-agent root gets a new invoke_workflow parent span.

Escape hatch: ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN=otel_semconv_1_36 (or 1) restores the previous format exactly. It is documented in the telemetry package doc and supported until at least March 2027. No Go API changes.

Testing Plan

Unit Tests:

  • All unit tests pass locally: the AGENTS.md module loop (go mod tidy -diff, go build, go test -race -shuffle=on, golangci-lint run v2.3.1) is clean in both modules.

Legacy is unchanged: no otel_semconv_1_36 golden changes in this PR; they are the ones #1640 records against main.

With your source change reverted and your tests kept, which test fails?
I mutated each new guard and branch on its own (30 mutations in this revision). All but one turn the suite red. The survivor is the span.IsRecording() check before converting tool arguments: it only skips work on a span the sampler dropped, and changes no output. Most are caught by TestTelemetrySchema/<scenario>/<id>. The rest have focused tests:

  • TestInstrumentInvocation_FirstErrorAndEarlyStop
  • TestInstrumentGenerateContent_EarlyStop
  • TestInferenceEventDescribesTheRequestTheSpanDoes: a model appending to req.Contents mid-call
  • TestStartExecuteToolSpan_ArgumentsAfterTheSamplingDecision
  • TestOversizedAttributeIsDropped: both schemas
  • TestEventContentAttributes
  • TestUseLegacySchema
  • TestTraceMergedToolCallsResult
  • TestExecuteTool_OmitsWhatIsUnknown
  • TestInvokeWorkflowNested
  • TestLogInferenceOperationDetails_ProviderName
  • TestTruthyValueDoesNotPutContentOnSpans
  • TestInferenceDetailsPartsMatchWebUI
  • TestUnrepresentableLogAttributesAreDropped

Manual End-to-End (E2E) Tests:
One real turn against gemini-2.5-flash on Vertex AI (llmagent plus one function tool), exported through telemetry.New(WithOtelToCloud(true)) to Cloud Trace and read back through the Cloud Trace API:

  • Default, SPAN_AND_EVENT: the tree is invoke_workflow → invoke_agent → generate_content / execute_tool / generate_content.
    • gen_ai.input.messages / output.messages are SLICE values and gen_ai.tool.call.arguments / result are MAP values.
    • Cloud Trace ingests them and shows them as JSON labels.
  • ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN=1 (the alias): today's tree, with JSON-string content.
  • Size probe: a span with a 100 KiB or 90 KiB structured attribute makes telemetry.googleapis.com answer 400 'Span.Attributes[1].value' is too large; at most 64.0K is allowed, rejecting the whole request. This is why the span guard stays in v2.

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code, plus fresh-context review passes; their findings are addressed.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules.

Additional context

Follow-ups, not in this PR:

@RKest
RKest force-pushed the feat/telemetry-semconv-default branch 2 times, most recently from 46a49d4 to 01f76b0 Compare September 24, 2026 15:04
@RKest RKest changed the title feat(telemetry)!: default to experimental GenAI semconv, with ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN=1 rollback feat(telemetry)!: default to experimental GenAI semconv, with ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN=otel_semconv_1_36 rollback Sep 24, 2026

@krisztianfekete krisztianfekete left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @RKest again 👋
We are also improving our OTel support as per kagent-dev/kagent#2907 over at kagent, and stumbled upon this PR. Since I am already deep into checking latest upstream genai semconv changes, I Ieft some comments/questions on the PR in case they are helpful. I see it's still a draft, but might help still as there are lots of changes.

Comment on lines +32 to +34
if useLegacySchema() || isWorkflowAgent(root) {
return run(ctx)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As per semconv invoke_workflow "SHOULD NOT be reported for standalone agent invocations", and its ADK example is Runner.run(...) "with multi-agent or graph workflow".

Could this span only be opened when the root has sub-agents, so a single LLM agent root stays invoke_agent?

Comment on lines +147 to +149
if ctx.Value(workflowScopeKey{}) != nil && !useLegacySchema() {
attrs = append(attrs, genAIWorkflowNested.Bool(true))
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gen_ai.workflow.nested isn't in the GenAI registry, and the spec says invoke_workflow SHOULD NOT be reported at all when it is an internal detail such as agent-as-tool. Wdyt about not opening the span in that case, instead of marking it? Or else propose the attribute upstream in semantic-conventions-genai together with adk-python, rather than defining it in the gen_ai.* namespace?

Comment on lines +215 to +222
// The event is emitted whatever the capture mode, so usage and finish reasons
// reach logs; messages are added only when content capture on events is on.
// No-op under the legacy schema, which logs through [logRequest] and
// [logResponse] instead.
func logInferenceOperationDetails(ctx context.Context, params GenerateContentParams, result generateContentResult) {
if useLegacySchema() {
return
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is opt-in? Can it be emitted only under EVENT_ONLY or SPAN_AND_EVENT?

// startGenerateContentSpan starts a new semconv generate_content span.
func startGenerateContentSpan(ctx context.Context, params GenerateContentParams) (context.Context, trace.Span) {
modelName := params.ModelName
attrs := []attribute.KeyValue{

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gen_ai.provider.name is Required on inference spans, but here it is only added to the event. Can it go on this span too? And can it come from an optional interface that any model.LLM can implement (for example ProviderName() string)? Otherwise non-Google models (OpenAI, Anthropic, Bedrock and others, as used by downstream projects like kagent) can never report it.

Comment on lines +228 to +230
if sys, ok := GenAISystemAttr(params.Backend); ok {
attrs = append(attrs, genAIProviderName.String(sys.Value.AsString()))
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This comes from the same GenAISystemAttr mapping, which only knows the Gemini and Vertex backends. For any other model.LLM the event has no gen_ai.provider.name, which the spec lists as Required on this event. Could this use the same provider source suggested for the span?

Comment on lines 257 to 263
func recordErrorAndStatus(span trace.Span, err error) {
if err == nil {
return
}
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

error.type is Conditionally Required on every GenAI span and on the inference event, but nothing sets it. Can this helper add it, using a bounded value such as the API or HTTP status code, the Go error type, or _OTHER? Generally, error rate queries and span-derived metrics rely on this.

)

const (
systemName = "gcp.vertex.agent"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The default now emits names that don't exist in 1.36, but the schema URL stays at v1.36.0; the same applies to the logger too. Since this PR is already the breaking change, can the schema URL follow the selected schema (so 1.36 only for legacy)?

…EMETRY_SCHEMA_VERSION_OPT_IN=otel_semconv_1_36 rollback

ADK now emits the experimental OpenTelemetry GenAI semantic conventions by
default, following adk-python's schema v2: one
gen_ai.client.inference.operation.details event per model call,
gen_ai.tool.call.arguments/result on execute_tool gated on span content
capture, and an entrypoint invoke_workflow span. Message content is
recorded as structured attribute values on spans as well as events, where
adk-python records JSON strings on spans.

ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN=otel_semconv_1_36 restores the previous
format, and otel_semconv_1_44 names the default. Values are
case-insensitive, and adk-python's 1 and 2 are accepted for them. No
otel_semconv_1_36 golden changes in this commit.

A model call and a runner invocation are traced only through
InstrumentGenerateContent and InstrumentInvocation, so the generate_content
span and its log records cannot drift apart.

Closes google#1634
@RKest
RKest force-pushed the feat/telemetry-semconv-default branch from 6cfc47a to 0174c33 Compare October 2, 2026 14:07
@RKest
RKest marked this pull request as ready for review October 2, 2026 14:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

telemetry: default to the experimental OpenTelemetry GenAI semantic conventions, with a one-env rollback

2 participants