test(telemetry): replace hand-written telemetry expectations with record/replay goldens - #1640
Conversation
41746ab to
fd20f6f
Compare
fd20f6f to
f802979
Compare
wolo-lab
left a comment
There was a problem hiding this comment.
One thing to fix before merge, and it is about the tests rather than the code: the new check that fails a recording when a log record is emitted outside any span has no test that exercises it.
Before this change the digest helper dropped such records silently by design (span_digest.go at the merge base, lines 102-109). Now BuildDigest fails instead, and the description lists that as one of the suite's guarantees. Replacing that t.Fatalf with continue leaves the whole suite green, and internal/telemetry/telemetrytest has no test file. So a later edit could quietly bring back the old drop. The check does work: a throwaway test that emits a record with no matching span hits the fatal. One thing to plan for when adding the test is that BuildDigest takes a concrete *testing.T, so asserting that it fails needs an error return or a testing.TB seam.
…ord/replay goldens
The functional telemetry tests compared each run against Go literals in
telemetrytestcase, so every telemetry change meant hand-editing expected
trees, and only one capture mode was covered. They now record the span
tree of a run, with its log records nested under the span they were
emitted in, as a golden, and replay it exactly. The digest follows
adk-python's _digests.py, except that a JSON-string attribute is wrapped
as {"JSON_STRING": ...}, so it stays distinct from a structured one.
Scenarios are typed: each has a package under scenarios/testcasedata
holding only its data (its own failure modes, and a Matrix of every
variant under every content capture mode and schema version), and
scenarios/testcaseimpl turns one into the agents, models and tools it
runs.
Every case is recorded under ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN set to
otel_semconv_1_36 and to otel_semconv_1_44. ADK does not read the variable
yet, so each pair is identical; the change that starts reading it shows
exactly which telemetry changes, and that the legacy format does not.
Re-record with: go generate ./internal/telemetry/functionaltest
Reverted the old behavior thanks |
f802979 to
df9d093
Compare
…ord/replay goldens (google#1640) The functional telemetry tests compared each run against Go literals in telemetrytestcase, so every telemetry change meant hand-editing expected trees, and only one capture mode was covered. They now record the span tree of a run, with its log records nested under the span they were emitted in, as a golden, and replay it exactly. The digest follows adk-python's _digests.py, except that a JSON-string attribute is wrapped as {"JSON_STRING": ...}, so it stays distinct from a structured one. Scenarios are typed: each has a package under scenarios/testcasedata holding only its data (its own failure modes, and a Matrix of every variant under every content capture mode and schema version), and scenarios/testcaseimpl turns one into the agents, models and tools it runs. Every case is recorded under ADK_TELEMETRY_SCHEMA_VERSION_OPT_IN set to otel_semconv_1_36 and to otel_semconv_1_44. ADK does not read the variable yet, so each pair is identical; the change that starts reading it shows exactly which telemetry changes, and that the legacy format does not. Re-record with: go generate ./internal/telemetry/functionaltest Co-authored-by: wolo <wolo@google.com>
Please ensure you have read the contribution guide before creating a pull request.
Linked issue
Problem:
The functional telemetry tests compared each run against hand-written Go literals in
internal/telemetry/telemetrytestcase. Only one content-capture mode was covered, and every telemetry change meant hand-editing expected span trees. The expectations could also not be compared with adk-python's recordings.Solution:
Record/replay goldens, following adk-python's
test_functional.py:internal/telemetry/functionaltest/testdata/<test_id>.json, where the id is<scenario>/[<variant>-]<capture>-<schema>, holds the span tree, with each log record nested under the span it was emitted in. The digest shape (root_span/name/attributes/status/children/logs, plus thePRESENTsentinel) follows adk-python's_digests.py, with one difference: a JSON-string attribute is wrapped as{"JSON_STRING": …}, so it stays distinct from a structured one.cmp.Diffbetween the recording and the golden. Recording fails on a log record emitted outside any span, and a golden that no case records fails the suite.go generate ./internal/telemetry/functionaltest.Scenarios are typed, and data is kept apart from setup:
Each scenario package has its own strongly typed
Failure(or none) and aMatrix()covering every variant under:no_content,span_only,event_only,span_and_event;ADK_TELEMETRY_SCHEMA_VERSION_OPT_INvalues:otel_semconv_1_36andotel_semconv_1_44.That comes to 80 goldens.
maindoes not read the schema variable yet, so eachotel_semconv_1_36/otel_semconv_1_44pair is identical here. #1635, which starts reading it, then changes onlyotel_semconv_1_44goldens, which shows exactly what telemetry changes and that the legacy format does not.Behavior change
Nothing. This PR changes tests only.
Testing Plan
Unit Tests:
go mod tidy -diff,go build,go test -race -shuffle=on,golangci-lint runv2.3.1) is clean in both modules.Coverage kept: each deleted expectation has a golden that pins at least as much:
AgentWithToolCase→toolcall/no_content-otel_semconv_1_36.jsonAgentWithToolCaptureContentCase→toolcall/span_and_event-otel_semconv_1_36.jsonWorkflow*Cases →dynamicworkflow/*no_content-otel_semconv_1_36.jsonThe goldens pin the legacy tool arguments and responses by value; the old expectations had
PRESENTthere.Determinism: 5 re-recordings under
-shuffle=onproduced byte-identical goldens.With your source change reverted and your tests kept, which test fails?
Not applicable: there is no source change. The goldens pin span names, attributes, status codes, the tree shape, and log event names, bodies and attributes. Like adk-python's digest, they do not pin span events or status descriptions.
Checklist