Summary
With the official Langfuse exporter:
install(OpenTelemetry) {
addLangfuseExporter(...)
setVerbose(true)
}
Langfuse Preview Output for a reasoning-model generation is the model's thinking trace, not the visible assistant reply.
Example: a classifier that must emit one token-word (RELEVANT / ESCALATE / OUT) shows a long English paraphrase of the user message in Preview. The generation's completion usage is 3 tokens — which matches RELEVANT, not the paragraph.
The agent itself is correct. Message.textContent() already reads only MessagePart.Text, so routing / downstream logic sees the real reply.
Root cause
LangfuseSpanAdapter.applyCompletionAttributes (same branch in applyPromptAttributes for assistant history) prefers reasoning over text when writing the OpenLLMetry attribute that Langfuse Preview maps to observation output:
when {
toolCalls.isNotEmpty() -> /* tool JSON */
reasoningParts.isNotEmpty() -> /* thinking, written as gen_ai.completion.{i}.content */
else -> /* actual MessagePart.Text */
}
Langfuse maps gen_ai.completion* → observation output. So thinking occupies the Preview slot and the real reply is dropped from that attribute. finish_reason is also omitted on the reasoning branch.
Providers such as xAI put thinking in reasoning_content → MessagePart.Reasoning and the short final answer in content → MessagePart.Text. xAI also reports completion_tokens without reasoning_tokens, which is why usage and Preview disagree.
gen_ai.output.messages already encodes both part types (type=text and type=reasoning). Langfuse does not render that attribute in Preview.
Reproduced against Koog 1.1.1; the same when is still on develop.
Proposed fix
- Keep
gen_ai.completion.{i}.content / gen_ai.prompt.{i}.content as tool calls or MessagePart.Text.
- Emit thinking on a sibling attribute (
gen_ai.completion.{i}.reasoning / gen_ai.prompt.{i}.reasoning).
- Keep
finish_reason on text completions even when reasoning parts are present.
Fix: #2215
Summary
With the official Langfuse exporter:
Langfuse Preview Output for a reasoning-model generation is the model's thinking trace, not the visible assistant reply.
Example: a classifier that must emit one token-word (
RELEVANT/ESCALATE/OUT) shows a long English paraphrase of the user message in Preview. The generation's completion usage is 3 tokens — which matchesRELEVANT, not the paragraph.The agent itself is correct.
Message.textContent()already reads onlyMessagePart.Text, so routing / downstream logic sees the real reply.Root cause
LangfuseSpanAdapter.applyCompletionAttributes(same branch inapplyPromptAttributesfor assistant history) prefers reasoning over text when writing the OpenLLMetry attribute that Langfuse Preview maps to observation output:Langfuse maps
gen_ai.completion*→ observationoutput. So thinking occupies the Preview slot and the real reply is dropped from that attribute.finish_reasonis also omitted on the reasoning branch.Providers such as xAI put thinking in
reasoning_content→MessagePart.Reasoningand the short final answer incontent→MessagePart.Text. xAI also reportscompletion_tokenswithoutreasoning_tokens, which is why usage and Preview disagree.gen_ai.output.messagesalready encodes both part types (type=textandtype=reasoning). Langfuse does not render that attribute in Preview.Reproduced against Koog 1.1.1; the same
whenis still ondevelop.Proposed fix
gen_ai.completion.{i}.content/gen_ai.prompt.{i}.contentas tool calls orMessagePart.Text.gen_ai.completion.{i}.reasoning/gen_ai.prompt.{i}.reasoning).finish_reasonon text completions even when reasoning parts are present.Fix: #2215