Repository navigation
fix(agents): keep Langfuse completion content as assistant text - #2215
Ordinary (OrdinarySF) wants to merge 1 commit into
Conversation
LangfuseSpanAdapter preferred MessagePart.Reasoning over Text when
writing gen_ai.completion.{i}.content. Langfuse Preview maps that
attribute to observation output, so reasoning traces (for example an
English paraphrase of a Chinese prompt) replaced the real reply such
as RELEVANT. Token usage still counted only the visible completion.
Write text or tool calls to content, emit reasoning on
gen_ai.completion.{i}.reasoning, and keep finish_reason when thinking
is present. Apply the same split to prompt history attributes.
Rebased onto develop after JetBrains#2248. Empty reasoning parts stay filtered
out. Non-empty reasoning is written to gen_ai.*.reasoning, not .content.
8062fb7 to
08cdd64
Compare
|
Rebased onto #2248 filters out a reasoning part whose content is empty (Gemini signature-only parts). That filter stays, on both prompt and completion attributes. This PR covers the remaining case: a reasoning part that has text. Langfuse Preview reads
Oleksandr Shylenko (@shilenkoalexander) #2248 landed in this same |
What
LangfuseSpanAdapterno longer writesMessagePart.Reasoningintogen_ai.completion.{i}.content(or the matching prompt attribute) when a visible assistant reply exists.gen_ai.completion.{i}.content/gen_ai.prompt.{i}.contentstay tool calls orMessagePart.Textgen_ai.completion.{i}.reasoning/gen_ai.prompt.{i}.reasoningfinish_reasonis kept on text completions even when reasoning parts are presentgen_ai.output.messagesis unchanged and still contains both part types.Why
Langfuse Preview maps
gen_ai.completion*to observation output. The adapter preferred reasoning over text:Reasoning models (Grok, o-series, DeepSeek, …) return thinking as
MessagePart.Reasoningand the short final answer asMessagePart.Text. Preview then showed the thinking (often an English paraphrase of a non-English prompt) whilegen_ai.usage.output_tokensstill counted only the visible completion — for example 3 tokens forRELEVANTnext to a long translation.The agent itself is fine:
Message.textContent()already reads onlyMessagePart.Text. This is an export mapping bug, not a model or Langfuse UI bug.How it was verified
LangfuseSpanAdapterTest:testCompletionAttributesPreferTextOverReasoningtestCompletionAttributesReasoningOnlyLeavesContentEmptytestPromptAttributesPreferTextOverReasoning./gradlew :agents:agents-features:agents-features-opentelemetry:jvmTestpasses locally.Notes
WeaveSpanAdapterstill prefers reasoning over text in the samewhen. Happy to follow up if you want the same split there.closes #2216